In this lesson
Phase 4 · Lesson 4.4

Transformers and Attention

The architecture behind ChatGPT, BERT, and every modern language model. Understand self-attention, and you understand the engine driving the AI revolution.

🕑 35 min read 📊 3 visualisations 💻 Hugging Face

Before Transformers: The Sequence Problem

Text is fundamentally different from an image. An image has fixed spatial structure: pixels at position (x, y) are always there. A sentence is a sequence of variable length, and the relationship between words can span long distances. In "The trophy did not fit in the bag because it was too large," what does "it" refer to? The trophy, not the bag. Understanding this requires relating "it" to "trophy" across eight words of distance.

For years, the dominant approach to sequence data was Recurrent Neural Networks (RNNs). An RNN processes a sequence one element at a time, passing a "hidden state" from step to step. The hidden state at step 50 is supposed to remember what happened at step 1. The problem is it often does not. Information gets diluted as it passes through many sequential multiplications.

Long Short-Term Memory networks (LSTMs), introduced by Hochreiter and Schmidhuber in 1997, added gating mechanisms that helped retain important information over longer distances. They powered breakthrough results in machine translation and speech recognition through the 2010s. But LSTMs still had a fundamental bottleneck: the entire history of a sequence had to be compressed into a single fixed-size hidden state vector before generating each output token.

In 2015, Bahdanau, Cho, and Bengio introduced the first attention mechanism for sequence-to-sequence models, allowing the decoder to look directly at all encoder states rather than relying solely on the compressed hidden state. This was enormously effective for machine translation. It was also the seed of a more radical idea: what if the entire architecture was built around attention, with no recurrence at all?

Why Recurrence Is Slow

RNNs process sequences step by step: step 2 cannot begin until step 1 is complete. This sequential dependency means you cannot parallelise the computation across the length of the sequence. On modern GPUs and TPUs, which are designed for massively parallel computation, this is a serious limitation. Transformers replace sequential recurrence with parallel attention, processing all positions simultaneously, which is a large part of why they scaled so dramatically.

The Core Idea: Attention

The intuition behind attention is something humans do naturally when reading. When you read the sentence "The scientist published the paper after she reviewed it carefully," and you want to understand what "she" refers to, your brain does not re-read the entire sentence. It pulls out the relevant word "scientist" and connects them. You are paying attention to specific parts of the input based on relevance to the current task.

In a Transformer, every token in a sequence is given the ability to directly attend to every other token, computing a weighted combination that reflects how relevant each position is. If "she" strongly attends to "scientist," the representation of "she" will be heavily influenced by the representation of "scientist," correctly capturing the coreference relationship.

The mechanism uses a query-key-value framework drawn from information retrieval. Think of it like a soft database lookup:

Query (Q)

What the current token is "looking for." Each token generates a query vector that represents what kind of information it needs from the context.

Key (K)

What each token "offers." Every token generates a key vector that describes the kind of information it contains. A high Query-Key match means high relevance.

Value (V)

The actual content contributed. If a token's key matches the query well, its value vector contributes heavily to the output. The output is a weighted average of all values.

Q, K, and V are not separate inputs. They are all derived from the same token embeddings through three different learned linear projections. This means the model learns what to look for, what to advertise, and what to contribute, all from the same underlying representations.

Self-Attention Step by Step

Let us walk through the exact computation. Given a sequence of tokens, each represented as a vector of dimension d_model:

Step 1: Project into Q, K, V

Multiply each token embedding by three learned weight matrices (W_Q, W_K, W_V) to get three vectors per token: a query Q, a key K, and a value V. The dimensions of Q and K must match (call it d_k) since we will compute their dot product.

Step 2: Compute Attention Scores

For each token acting as a query, compute a dot product with every token's key. A high dot product means the query and key are aligned, that is, this pair of tokens has a strong relationship.

Step 3: Scale and Softmax

Divide each score by the square root of d_k. This scaling prevents the dot products from growing too large in magnitude (which would push the softmax into regions where gradients become tiny). Then apply softmax to get a probability distribution over all positions, the attention weights.

Attention(Q, K, V) = softmax( QKT / √dk ) · V
The scaled dot-product attention formula (Vaswani et al., 2017). QKT computes all pairwise scores. √dk prevents large magnitudes. Softmax normalises to weights. V provides the content.

Step 4: Weighted Sum of Values

Multiply the attention weights by the value vectors and sum. The output for each token is a weighted mixture of all value vectors, where the weights reflect how relevant each position was. Tokens that scored highly in the attention contribute more to the output.

Self-Attention Weight Matrix Row = query token. Column = key token. Darker = higher attention weight. "The" "cat" "is" "gone" KEY TOKENS (attending to) "The" "cat" "is" "gone" QUERY TOKENS (looking) 0.55 0.25 0.12 0.08 0.15 0.50 0.15 0.20 0.20 0.30 0.35 0.15 0.10 0.35 0.10 0.45 = 1.00 = 1.00 = 1.00 = 1.00 Low attention Medium High attention Illustrative weights. Each row sums to 1.0 after softmax.

An illustrative self-attention weight matrix for a 4-token sentence. "gone" attends most strongly to itself (0.45) and to "cat" (0.35), capturing the semantic relationship. Every row is a softmax distribution over all key positions.

A Concrete Implementation

The mathematics of scaled dot-product attention can be implemented in about ten lines of Python:

Python (NumPy)
import numpy as np

def softmax(x):
    # Subtract max for numerical stability — prevents overflow in exp()
    exp_x = np.exp(x - np.max(x, axis=-1, keepdims=True))
    return exp_x / exp_x.sum(axis=-1, keepdims=True)

def scaled_dot_product_attention(Q, K, V):
    """
    Q: (seq_len, d_k) — queries
    K: (seq_len, d_k) — keys
    V: (seq_len, d_v) — values
    """
    d_k = Q.shape[-1]

    # Step 1: compute all pairwise scores — (seq_len, seq_len)
    scores = Q @ K.T / np.sqrt(d_k)

    # Step 2: apply softmax row-wise to get attention weights
    weights = softmax(scores)

    # Step 3: weighted sum of value vectors
    output = weights @ V

    return output, weights

# Example: 4 tokens, d_k = 4
np.random.seed(42)
seq_len, d_k = 4, 4

Q = np.random.randn(seq_len, d_k)
K = np.random.randn(seq_len, d_k)
V = np.random.randn(seq_len, d_k)

output, attn_weights = scaled_dot_product_attention(Q, K, V)

print("Attention weights (each row sums to 1):")
print(attn_weights.round(3))
print(f"\nOutput shape: {output.shape}")
Attention weights (each row sums to 1): [[0.429 0.112 0.198 0.261] [0.277 0.271 0.319 0.133] [0.151 0.381 0.272 0.196] [0.087 0.340 0.402 0.171]] Output shape: (4, 4)

This is the complete core of self-attention. Each row in the weight matrix is a distribution over all four positions. The output for each token is a weighted blend of all value vectors according to these weights.

Multi-Head Attention

A single attention head learns one way to relate tokens. But a sentence has multiple types of relationships simultaneously. "John gave Mary the book" contains relationships between giver and receiver, between action and object, between subject and verb. A single attention pattern cannot capture all of these at once.

Multi-head attention runs several attention operations in parallel, each with its own independent W_Q, W_K, and W_V projection matrices. Each "head" is free to learn a different type of relationship. The outputs from all heads are concatenated and projected through a final linear layer.

The original Transformer used 8 heads with d_model = 512. Each head operates in a reduced dimension of 512 / 8 = 64, so the total computational cost is comparable to a single head at full dimension. Research on analysing what individual heads learn has found that some heads specialise in syntactic relationships (subject-verb agreement), some in coreference (pronoun-noun), and some in positional patterns, though this specialisation is not guaranteed or required by the design.

Multi-Head Attention in One Line

MultiHead(Q, K, V) = Concat(head_1, ..., head_h) W_O, where each head_i = Attention(Q W_Q_i, K W_K_i, V W_V_i). The W_O matrix projects the concatenated output back to d_model dimensions.

The Transformer Block

Self-attention alone is not a complete model. The full Transformer block wraps it with several additional components, each serving a specific engineering purpose.

Transformer Encoder Block Token Embeddings + Positional Encoding Multi-Head Self-Attention Q, K, V from same input sequence Add & LayerNorm residual connection Add & Layer Norm Feed-Forward Network Dense(d_ff, ReLU) → Dense(d_model) Applied identically to each position Add & Layer Norm Contextual Representations Stack N times (original: N = 6) This entire block stacks N times. Each block refines the representations.

One Transformer encoder block. The green dashed lines are residual (skip) connections, exactly as in ResNet. They help gradients flow through deep stacks. Layer Norm stabilises training. The feed-forward network processes each position independently after attention has aggregated context.

Residual Connections

The output of each sub-layer is added to its input before normalisation. This is the same skip connection idea from ResNet (He et al., 2015), and it is critical for training deep stacks of Transformer blocks without vanishing gradients.

Layer Normalisation

Introduced by Ba, Kiros, and Hinton in 2016, Layer Norm normalises across the feature dimension for each token independently. Unlike Batch Norm, it does not depend on batch size, making it better suited to variable-length sequences.

Feed-Forward Network

After attention aggregates context, a two-layer MLP (Dense → ReLU → Dense) processes each token independently. This is where much of the model's "knowledge" is stored. In the original paper, the inner dimension d_ff = 2,048 (four times d_model = 512).

Positional Encoding

Attention has no inherent notion of order. To tell the model where each token sits in the sequence, positional information is added to the embeddings. The original paper used fixed sinusoidal functions. Modern models (GPT, BERT variants) typically use learned positional embeddings.

The 2017 Breakthrough: "Attention Is All You Need"

The Transformer architecture was introduced in the paper "Attention Is All You Need" by Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, all researchers at Google at the time. It was presented at NeurIPS 2017.

The paper's central claim was bold: recurrence and convolution were not necessary for sequence modelling. Attention alone, applied in parallel across all positions, was sufficient and indeed superior. The model they proposed achieved state-of-the-art results on English-to-German and English-to-French machine translation tasks, while training far faster than the recurrent models of the time.

The base Transformer had these specs:

# Original Transformer "base" hyperparameters (Vaswani et al., 2017)
d_model    = 512   # embedding / model dimension
d_ff       = 2048  # feed-forward inner dimension (4× d_model)
num_heads  = 8     # attention heads per block
d_k = d_v  = 64   # dimension per head (512 / 8)
num_layers = 6     # number of encoder blocks (and 6 decoder blocks)
dropout    = 0.1   # dropout rate applied throughout
# Total parameters: approximately 65 million

The paper also introduced the concept of scaled dot-product attention. Removing the scaling factor 1/sqrt(d_k) causes the dot products to grow large in high dimensions, pushing softmax into saturation regions where gradients nearly vanish. The scaling was a practical and important engineering decision.

Why the Paper Name Matters

The title is a direct challenge to the field's assumptions. Attention mechanisms had been used as add-ons to RNNs. The claim that attention alone, without any recurrence, was not just sufficient but better was a paradigm shift. The paper has been cited well over 100,000 times and spawned an entire new era of AI architecture design.

BERT and GPT: Two Strategies, One Architecture

The Transformer has an encoder stack and a decoder stack. Different research groups discovered that pre-training just one of these components on massive text datasets, then fine-tuning on specific tasks, produced remarkably capable models. Two paradigms emerged.

BERT (Encoder-Only)

Bidirectional Encoder Representations from Transformers. Google, 2018 (Devlin, Chang, Lee, Toutanova).

  • Uses only the encoder stack
  • Each token attends to all tokens in both directions simultaneously (bidirectional)
  • Pre-trained with Masked Language Modelling: randomly mask 15% of tokens, predict them from context
  • Best for: understanding tasks — question answering, sentiment analysis, named entity recognition, text classification
  • Not designed to generate text

GPT (Decoder-Only)

Generative Pre-trained Transformer. OpenAI, 2018 (Radford, Narasimhan, Salimans, Sutskever).

  • Uses only the decoder self-attention with causal masking
  • Each token can only attend to previous tokens (left-to-right only)
  • Pre-trained with Causal Language Modelling: predict the next token given all previous tokens
  • Best for: generation tasks — writing, summarisation, conversation, code
  • ChatGPT, GPT-4, and most modern chatbots use this paradigm

The Scale-Up Story

2018
GPT-1 (OpenAI): 117 million parameters. Showed that language model pre-training transfers to many downstream tasks.
2018
BERT-base (Google): 110 million parameters. Set new records on 11 NLP benchmarks simultaneously. BERT-large: 340M parameters.
2019
GPT-2 (OpenAI): 1.5 billion parameters. Generated coherent multi-paragraph text; OpenAI initially withheld the largest version citing misuse concerns.
2020
GPT-3 (OpenAI): 175 billion parameters. Demonstrated few-shot learning: given just a few examples in the prompt, it could perform tasks it was never explicitly trained on.
2022
ChatGPT (OpenAI): GPT-3.5 fine-tuned with Reinforcement Learning from Human Feedback (RLHF). Reached 100 million users in two months, the fastest product adoption in history to that point.
2023
GPT-4 (OpenAI) and Claude (Anthropic): multimodal, more capable, more safety-aligned large language models built on the Transformer decoder paradigm.
Causal Masking Explained

In a decoder-only model, when predicting token 5, the model cannot look at tokens 6, 7, 8... (they do not exist yet in generation). This is enforced by a causal mask: a triangular mask that sets attention scores to negative infinity for all future positions before softmax, making their weights effectively zero. During training, the model processes all positions in parallel but each position only "sees" its left context, mimicking the sequential generation process efficiently.

Using Transformers Today: Hugging Face

Training a Transformer from scratch requires enormous data and compute. In practice, you almost never do this. Instead, you use pre-trained models through the Hugging Face transformers library, the dominant open-source platform for working with Transformer models. It provides a unified API for hundreds of pre-trained models in dozens of languages.

Install First

Run this once in your terminal or a Colab cell: pip install transformers torch. The library handles downloading model weights automatically on first use.

Python (Hugging Face Transformers)
from transformers import pipeline

# ── 1. Sentiment Analysis using a fine-tuned BERT model ───────────────
# Default model: distilbert-base-uncased-finetuned-sst-2-english
# (a compressed BERT fine-tuned on Stanford Sentiment Treebank)
classifier = pipeline("sentiment-analysis")

texts = [
    "I really enjoyed this course, the explanations are excellent!",
    "This was confusing and poorly structured.",
    "Convolutional neural networks are used in image recognition."
]

results = classifier(texts)
for text, result in zip(texts, results):
    print(f"Text: '{text[:50]}...'")
    print(f"  Label: {result['label']}, Score: {result['score']:.4f}\n")

# ── 2. Text Generation using GPT-2 ────────────────────────────────────
generator = pipeline("text-generation", model="gpt2")

prompt = "Artificial intelligence is transforming"
output = generator(
    prompt,
    max_new_tokens=50,       # generate 50 new tokens after the prompt
    num_return_sequences=1,
    do_sample=True,          # sample from the distribution (not greedy)
    temperature=0.8          # lower = more predictable, higher = more creative
)

print("Generated text:")
print(output[0]['generated_text'])

# ── 3. Question Answering using BERT ──────────────────────────────────
qa = pipeline("question-answering")

context = """
The Transformer architecture was introduced in 2017 by Vaswani et al.
in the paper 'Attention Is All You Need'. It relies entirely on
self-attention mechanisms and has become the foundation of modern
language models including BERT and GPT.
"""

result = qa(question="When was the Transformer introduced?", context=context)
print(f"\nQuestion: When was the Transformer introduced?")
print(f"Answer: {result['answer']} (confidence: {result['score']:.4f})")
Text: 'I really enjoyed this course, the explanations are ...' Label: POSITIVE, Score: 0.9997 Text: 'This was confusing and poorly structured....' Label: NEGATIVE, Score: 0.9995 Text: 'Convolutional neural networks are used in image rec...' Label: POSITIVE, Score: 0.9748 Generated text: Artificial intelligence is transforming the way businesses operate, enabling companies to automate complex tasks, analyse vast amounts of data, and make decisions with greater accuracy and speed than was previously possible. Question: When was the Transformer introduced? Answer: 2017 (confidence: 0.9961)

What Just Happened?

The pipeline function downloaded a pre-trained model (several hundred MB), loaded its weights, and ran inference. For sentiment analysis, it used a DistilBERT model fine-tuned on human-labelled movie and review data. For question answering, it used a BERT model trained to extract spans from context paragraphs. None of the heavy lifting required any data or training from us.

The Power of Pre-Training

These models were pre-trained on billions of words of text, which gave them broad language understanding. They were then fine-tuned on labelled examples specific to each task. This pre-train-then-fine-tune paradigm is covered in depth in Lesson 4.5 on Transfer Learning. It is the reason a beginner can achieve expert-level NLP results with three lines of code.

Key Takeaways

Coming Up: Transfer Learning and Fine-Tuning

In Lesson 4.5, you will see how to take a pre-trained model and adapt it to your own specific task with a small amount of data. This technique, called fine-tuning, is how most real-world AI applications are built today. You do not need to train a model from scratch. You start from a powerful foundation and teach it your specific domain.

Practice Notebook
Run this lesson's code in Google Colab
All examples + challenges · Free GPU included · No setup needed
Open In Colab
Progress
Done with this lesson?
Mark it complete to track your progress.